NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

MATCH: Metadata-Aware Text Classification in A Large Hierarchy

https://doi.org/10.1145/3442381.3449979

Zhang, Yu; Shen, Zhihong; Dong, Yuxiao; Wang, Kuansan; Han, Jiawei (April 2021, WWW '21: The Web Conference 2021)
null (Ed.)
Multi-label text classification refers to the problem of assigning each given document its most relevant labels from a label set. Commonly, the metadata of the given documents and the hierarchy of the labels are available in real-world applications. However, most existing studies focus on only modeling the text information, with a few attempts to utilize either metadata or hierarchy signals, but not both of them. In this paper, we bridge the gap by formalizing the problem of metadata-aware text classification in a large label hierarchy (e.g., with tens of thousands of labels). To address this problem, we present the MATCH solution—an end-to-end framework that leverages both metadata and hierarchy information. To incorporate metadata, we pre-train the embeddings of text and metadata in the same space and also leverage the fully-connected attentions to capture the interrelations between them. To leverage the label hierarchy, we propose different ways to regularize the parameters and output probability of each child label by its parents. Extensive experiments on two massive text datasets with large-scale label hierarchies demonstrate the effectiveness of MATCH over the state-of-the-art deep learning baselines.
more » « less
Full Text Available
Heterogeneous Graph Transformer

https://doi.org/10.1145/3366423.3380027

Hu, Ziniu; Dong, Yuxiao; Wang, Kuansan; Sun, Yizhou (January 2020, WWW '20: Proceedings of The Web Conference 2020)
null (Ed.)
Full Text Available
TaxoExpan: Self-supervised Taxonomy Expansion with Position-Enhanced Graph Neural Network

https://doi.org/10.1145/3366423.3380132

Shen, Jiaming; Shen, Zhihong; Xiong, Chenyan; Wang, Chi; Wang, Kuansan; Han, Jiawei (April 2020, WWW '20: The Web Conference 2020)

Taxonomies consist of machine-interpretable semantics and provide valuable knowledge for many web applications. For example, online retailers (e.g., Amazon and eBay) use taxonomies for product recommendation, and web search engines (e.g., Google and Bing) leverage taxonomies to enhance query understanding. Enormous efforts have been made on constructing taxonomies eithermanually or semi-automatically. However, with the fast-growing volume of web content, existing taxonomies will become outdated and fail to capture emerging knowledge. Therefore, in many applications, dynamic expansions of an existing taxonomy are in great demand. In this paper, we study how to expand an existing taxonomy by adding a set of new concepts. We propose a novel self-supervised framework, named TaxoExpan, which automatically generates a set of ⟨query concept, anchor concept⟩ pairs from the existing taxonomy as training data. Using such self-supervision data, TaxoExpan learns a model to predict whether a query concept is the direct hyponym of an anchor concept. We develop two innovative techniques in TaxoExpan: (1) a position-enhanced graph neural network that encodes the local structure of an anchor concept in the existing taxonomy, and (2) a noise-robust training objective that enables the learned model to be insensitive to the label noise in the self-supervision data. Extensive experiments on three large-scale datasets from different domains demonstrate both the effectiveness and the efficiency of TaxoExpan for taxonomy expansion.
more » « less
Full Text Available
GPT-GNN: Generative Pre-Training of Graph Neural Networks

https://doi.org/10.1145/3394486.3403237

Hu, Ziniu; Dong, Yuxiao; Wang, Kuansan; Chang, Kai-Wei; Sun, Yizhou (January 2020, Proc. of 2020 ACM SIGKDD Int. Conf. on Knowledge Discovery and Data Mining (KDD’20))

Full Text Available

Search for: All records